Papers with inference-time interventions
Steering Large Language Models for Machine Translation Personalization (2026.eacl-long)
Copied to clipboard
| Challenge: | Recent advances in interpretability research have highlighted the effectiveness of steering methods for MT personalization. |
| Approach: | They examine steering strategies for personalizing automatic translations when few examples are available. |
| Outcome: | The proposed steering methods yield higher inference-time computational efficiency than prompting approaches. |
Probing Political Ideology in Large Language Models: How Latent Political Representations Generalize Across Tasks (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models encode rich internal representations of political ideology, but it remains unclear how these representations contribute to model decision-making. |
| Approach: | They apply inference-time interventions to steer a decoder-only transformer along learned ideological directions . they find that learned ideological representations generalize well to bias detection, but not as well to voting simulations . |
| Outcome: | The proposed model steers a transformer along learned ideological directions . political bias detection, voting preference simulation and bias neutralization are tested . |